Papers with spoken language processing
Sounding Board: A User-Centric and Content-Driven Social Chatbot (N18-5)
Copied to clipboard
Hao Fang, Hao Cheng, Maarten Sap, Elizabeth Clark, Ari Holtzman, Yejin Choi, Noah A. Smith, Mari Ostendorf
| Challenge: | Sounding Board is a social chatbot that can hold a coherent conversation with humans . the system is user-centric in that users can control the topic of conversation, while the system adapts to the user's needs. |
| Approach: | They present Sounding Board, a social chatbot that won the 2017 Amazon Alexa Prize. |
| Outcome: | The system is user-centric in that users can control the topic of conversation, while the system adapts to the user's needs. |
CPJD Corpus: Crowdsourced Parallel Speech Corpus of Japanese Dialects (L18-1)
Copied to clipboard
| Challenge: | Various corpora of dialects have been collected using a well-equipped recording environment due to geographical and expense issues. |
| Approach: | They construct a crowdsourced parallel speech corpus of Japanese dialects using crowdsourcing platforms. |
| Outcome: | The proposed corpus includes parallel text and speech data of 21 Japanese dialects. |
MYCanCor: A Video Corpus of spoken Malaysian Cantonese (L18-1)
Copied to clipboard
| Challenge: | The corpus consists of 20 hours of video recordings of spontaneous talk-in-interaction typically involving 2-4 speakers. |
| Approach: | the corpus consists of 20 hours of video recordings of spontaneous talk-in-interaction typically involving 2-4 speakers. |
| Outcome: | the corpus consists of 20 hours of video recordings of spontaneous talk-in-interaction typically involving 2-4 speakers. |
SMASH Corpus: A Spontaneous Speech Corpus Recording Third-person Audio Commentaries on Gameplay (2020.lrec-1)
Copied to clipboard
| Challenge: | Developing a spontaneous speech corpus is important for spoken language research . a corpus of spontaneous speech is needed to develop these techniques . |
| Approach: | They propose to use Japanese male commentators' spontaneous speech to construct a SMASH corpus . they use transcriptions and topic tags to annotate the commentaries and report some results . |
| Outcome: | The proposed corpus includes spontaneous speech of two Japanese male commentators . the authors report that the annotations yielded a better corpus than the previous methods . |
POWSM: A Phonetic Open Whisper-Style Speech Foundation Model (2026.acl-long)
Copied to clipboard
Chin-Jou Li, Kalvin Chang, Shikhar Bharadwaj, Eunjung Yeo, Kwanghee Choi, Jian Zhu, David R. Mortensen, Shinji Watanabe
| Challenge: | Phone-level modeling of speech is a common approach to speech recognition, but it relies on task-specific architectures and datasets. |
| Approach: | They propose a phonetic framework capable of performing multiple phone-related tasks . they propose 'Phonetic Open Whisper-style Speech Model' that can perform these tasks together . |
| Outcome: | The proposed model outperforms or matches specialized PR models of similar size while supporting G2P, P2G, and ASR. |